Restore VMware-to-KVM migration changes from #13656 - #14256
andrijapanicsb wants to merge 5 commits into
Conversation
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## main #14256 +/- ##
============================================
+ Coverage 19.91% 20.08% +0.17%
- Complexity 20194 20693 +499
============================================
Files 6373 6427 +54
Lines 577230 584934 +7704
Branches 70696 71665 +969
============================================
+ Hits 114942 117476 +2534
- Misses 449722 454592 +4870
- Partials 12566 12866 +300
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
|
Packages build from the identical code: Packaging results:
Test packages are available at: |
|
ShapeBlue clean/successfull packaging pass - link: #13656 (comment) |
|
ShapeBlue BlueOrangutan (Marvin tests - all 156 passed with zero failures) - link: #13656 (comment) |
|
@alexandremattioli cant see you from the dropdown in "reviewers" so just pinging you this way, especially if you have any capacity for testing etc. |
|
@ACSHomeBot package G |
Packaging results:
Test packages are available at:
The packages previously published for |
|
Building (as you can see above) new packages for this one @DaanHoogland - just to have it officially/fresh |
|
@blueorangutan package |
1 similar comment
|
@blueorangutan package |
|
@winterhazel thx for the review request |
| cycle.setState(VmwareCbtMigrationCycle.State.Created); | ||
| cycle.setDescription("Creating final VMware CBT snapshot for cutover"); | ||
| cycle.setUpdated(new Date()); | ||
| cycle = vmwareCbtMigrationCycleDao.persist(cycle); |
There was a problem hiding this comment.
This DAO call could also be included within the try block for added safety and consistent error handling. Since the methods below already have their DAO interactions protected by the same try-catch construct, it may be beneficial to keep this call within that scope as well to handle any unexpected failure scenarios gracefully.
There was a problem hiding this comment.
You are right that this insert sits outside the catch. Simply moving it into the try block would not be safe because the catch path assumes the cycle row exists. I am handling the insert-failure case separately and adding a focused test.
|
@blueorangutan test |
|
@nvazquez a [SL] Trillian-Jenkins test job (ol8 mgmt + kvm-ol8) has been kicked to run smoke tests |
|
[SF] Trillian Build Failed (tid-17040) |
Signed-off-by: andrijapanicsb <andrija.panic@gmail.com>
Signed-off-by: andrijapanicsb <andrija.panic@gmail.com>
|
@ACSHomeBot package G |
Packaging results:
Test packages are available at:
The packages previously published for |
|
@blueorangutan package |
|
@andrijapanicsb a [SL] Jenkins job has been kicked to build packages. It will be bundled with no SystemVM templates. I'll keep you posted as I make progress. |
|
Packaging result [SF]: ✔️ el8 ✔️ el9 ✔️ el10 ✔️ debian ✔️ suse15. SL-JID 19349 |
|
@andrijapanicsb reviewing functionally in my labs |
Port the production fix from apache/cloudstack PR apache#14195, source commit 0cb24ba (merged into 4.22 as 273b2ea). Reuse the supplied template and load its details without a findById lookup that filters removed records. Preserve template CPU-detail inheritance and explicit VM overrides. Add six regression tests for dummy and active templates, non-VO templates, CPU override precedence and missing templates. Co-authored-by: Abhisar Sinha <63767682+abh1sar@users.noreply.github.com>
|
@ACSHomeBot package G |
Packaging results:
Test packages are available at: |
|
@blueorangutan package |
|
@andrijapanicsb a [SL] Jenkins job has been kicked to build packages. It will be bundled with no SystemVM templates. I'll keep you posted as I make progress. |
|
Packaging result [SF]: ✔️ el8 ✔️ el9 ✔️ el10 ✔️ debian ✔️ suse15. SL-JID 19364 |
|
@blueorangutan test |
|
@NuxRo a [SL] Trillian-Jenkins test job (ol8 mgmt + kvm-ol8) has been kicked to run smoke tests |
|
[SF] Trillian Build Failed (tid-17057) |
Skip target cleanup on cancel and delete when an imported VM is recorded, including imports that completed after cancellation. Re-read the migration before cleanup to protect VM references recorded since the caller loaded it. Keep source snapshot cleanup and migration record deletion independent. Add regression coverage for both cleanup entry points, all migration states, stale caller records, and record-only deletion with source cleanup. This guard does not resolve imports racing after the final read. Signed-off-by: andrijapanicsb <andrija.panic@gmail.com>
|
@ACSHomeBot package G |
|
Accepted: Planned: RPM + DEB, KVM SystemVM. Status: Queued for worker G. Jobs ahead: 0. |
|
Package building started on worker G. |
Packaging results:
Test packages are available at: |
Published package retention updateThe packages previously published for |
| // Import may finish after cancellation and record the VM without changing the Cancelled state. | ||
| // Its target disks now belong to that VM, so neither cancel nor delete may remove them. | ||
| // This does not prevent source snapshot cleanup or deletion of the migration record. |
There was a problem hiding this comment.
| // Import may finish after cancellation and record the VM without changing the Cancelled state. | |
| // Its target disks now belong to that VM, so neither cancel nor delete may remove them. | |
| // This does not prevent source snapshot cleanup or deletion of the migration record. |
| // This does not prevent source snapshot cleanup or deletion of the migration record. | ||
| Long importedVmId = migration.getVmId(); | ||
| if (importedVmId == null) { | ||
| // Re-read before cleanup: an import may have recorded its VM since this caller loaded the migration. |
There was a problem hiding this comment.
| // Re-read before cleanup: an import may have recorded its VM since this caller loaded the migration. |
rp-
left a comment
There was a problem hiding this comment.
Hey Andrija,
sorry but I did review again the Linstor parts and found 3 issues that should be probably fixed before merging.
| // A Linstor volume is a DRBD device that only appears on this host once its resource | ||
| // is made available here (a diskless assignment); connect it before qemu-img inspects | ||
| // the device and release the diskless assignment afterwards (the replicated data on | ||
| // the storage nodes is untouched). RBD needs no such step. | ||
| boolean linstorConnected = false; | ||
| if (StoragePoolType.Linstor.equals(pool.getType())) { | ||
| linstorConnected = storagePoolMgr.connectPhysicalDisk(pool.getType(), pool.getUuid(), volumePath, null); | ||
| } | ||
| try { | ||
| return addVolumeByVolumePath(command, storagePool, volumePath); | ||
| } finally { | ||
| if (linstorConnected) { | ||
| storagePoolMgr.disconnectPhysicalDisk(pool.getType(), pool.getUuid(), volumePath); | ||
| } | ||
| } |
There was a problem hiding this comment.
The connect step is only added for the single-path case. addAllVolumes (L162–174, not part of this diff) runs getDiskFileInfo on each disk from listPhysicalDisks. For Linstor, disk.getPath() is /dev/drbd/by-res/cs-…/0, which only exists if the resource has a replica or diskless resource on this host. For every other volume qemu-img info fails, info == null, and the continue skips it without any error. In the UI, the import-data-disk list for a Linstor pool therefore shows only the volumes that happen to be on whichever host the management server sent the command to.
The listing is also expensive on large controllers. For each resource there are 2 API calls in getPhysicalDisk (volumeDefinitionList and viewResources), 1 in getVolumeInUseNode (resourceList), and 2 qemu-img info runs.
For Linstor you should probably, don't inspect devices when listing. The format is always RAW, size and in-use state come from the controller, and backing files and qcow2 encryption can't apply to a raw DRBD device. A single viewResources call (filtered to the pool's resource group) can provide name, size and in-use state for all resources at once. Connect and inspect only in the single-path case, as now.
| return checkRbdVolume(command, pool, vol); | ||
| } finally { | ||
| if (linstorConnected) { | ||
| poolMgr.disconnectPhysicalDisk(storageFilerTO.getType(), storageFilerTO.getUuid(), srcFile); |
There was a problem hiding this comment.
LinstorStorageAdaptor.connectPhysicalDisk returns true whenever resourceMakeAvailableOnNode succeeds, including when the resource was already on this node. linstorConnected therefore means "the call succeeded", not "we created the local resource". The disconnect then goes through tryDisconnectLinstor, which:
- deletes the local resource if it is diskless and not a tiebreaker, even if it existed before this check (e.g. placed by an admin, or left over from another operation);
- calls removeTwoPrimariesProps if the resource is in use on another node. If that volume is being live-migrated, this strips allow-two-primaries and protocol in the middle of the migration.
I suggest: before connecting, check whether a resource for this name already exists on the local node (one viewResources filtered to the local node and this resource), and only disconnect if this check created it. If the in-use check reports another node, skip the disconnect completely so the allow-two-primaries properties are left alone. Same pattern applies to LibvirtGetVolumesOnStorageCommandWrapper.java L72–82 (see comment there).
| return addVolumeByVolumePath(command, storagePool, volumePath); | ||
| } finally { | ||
| if (linstorConnected) { | ||
| storagePoolMgr.disconnectPhysicalDisk(pool.getType(), pool.getUuid(), volumePath); |
There was a problem hiding this comment.
Same issue as in LibvirtCheckVolumeCommandWrapper L89: this disconnect can remove a local diskless resource that existed before the call, and can strip allow-two-primaries from a resource that is in use elsewhere. It should only disconnect if this call created the local resource.
| List<ResourceDefinition> rscDfns = LinstorUtil.getRDListStartingWith(api, LinstorUtil.RSC_PREFIX); | ||
| for (ResourceDefinition rscDfn : rscDfns) { | ||
| if (rscGroup != null && !rscGroup.equalsIgnoreCase(rscDfn.getResourceGroupName())) { | ||
| continue; | ||
| } | ||
| String name = rscDfn.getName().substring(LinstorUtil.RSC_PREFIX.length()); |
There was a problem hiding this comment.
This check keeps the listing to resources in the pool's own resource group. But importVolume path=… (GetVolumesOnStorage with a volume path) and importVm diskpath=… (CheckVolume) go straight to getPhysicalDisk(name), which accepts any cs-* resource on the controller. The duplicate checks on the management server only look within the target pool:
VolumeImportUnmanageManagerImpl.java:370:volumeDao.findByPoolIdAndPath(pool.getId(), volumePath)UnmanagedVMsManagerImpl.java:3237:findByPoolIdAndPath(poolId, diskPath)
Resource names are unique per LINSTOR controller, and several CloudStack pools on one controller (different resource groups, e.g. SSD and HDD) is a common setup. Scenario:
- Pool A (resource group rg-ssd) has volume X, i.e. resource cs-X, attached to a stopped VM.
- An admin runs importVolume with storageid= (resource group rg-hdd, same controller) and path=X.
- All checks pass. The in-use check doesn't help because nothing is running.
- Two CloudStack volumes now point at the same DRBD resource. Deleting either one calls
deleteResourceDefinition(cs-X), and the other volume's data is gone.
Suggested fix: on the management server, before sending CheckVolume or GetVolumesOnStorage for a Linstor pool, look up cs-<path> on the controller (the driver already builds the API client via LinstorUtil.getLinstorAPI(pool.getHostAddress(), …)) and reject the import if ResourceDefinition.getResourceGroupName() doesn't match the pool's resource group. That uses the controller's own data, covers resources this CloudStack database doesn't know about, and avoids comparing controller addresses across pools. It applies to both importVolume and importVm importsource=shared.
Description
This PR restores the exact content of #13656, which was merged as
0a5bf30af32bdea5a209f3f993cbd6a43301d0f9and then reverted by510d0ec3785efe3cce65ccd1247682b91f4492d0.The restoration commit
44a8713671d3a8830342762e88975ad3fd3426c7reproduced the original merge's Git tree (d3be509ec552975ff628e8f2ada6f4d46d1f109d) and stable patch ID (43eeb0cbbadf2e566bc43780ee1c5244888451d0).A subsequent focused commit,
a5578d6f53b51e2b67c7c537d3b7217c6d44eb3d, fixes the per-disk checkpoint used between warm CBT delta cycles. VMware'sDiskChangeInfodoes not contain a change ID; the new checkpoint is now read from the cycle snapshot's disk backing. A second focused commit,6fa7d30b47, handles failures when the final agent command throws or the final CBT cycle cannot be recorded. The original restoration is otherwise unchanged.A further focused commit,
2a3598ad6a, ports the upstream dummy-template import fix described below, with regression tests.The complete feature description and implementation details remain available in
the original PR:
#13656
The original change, before these focused fixes, was approved by two independent committers:
Existing upstream importVM regression
During the new QA run, VM import without a supplied template failed with
Unable to find template with id ... for virtual machine import. This is an existing upstream regression, not introduced by the VMware-to-KVM restoration in this PR.PR #12793, merged as
a01fb0be34b2774d8fb7853703b36364444398e4, added template-detail loading so imported VMs inherit settings such asguest.cpu.mode=host-passthrough. However, its additionalfindById(template.getId())lookup excludes soft-deleted records. The defaultVM Import Default Templateis deliberately stored in the removed state, so this lookup returns null and rejects an otherwise valid import.This was already fixed on the
4.22branch by PR #14195, using source commit0cb24ba8967e14e9577dde7754eeb4a93e6e23e7(merged as273b2ea32bfa24396431b9bb0df055e0ac6c1c02). That fix is still absent frommainas checked on 1 October 2026 at1a48a87587c03472feefb081cfac71b2ebd0f407, which retains the failing lookup.This PR ports the same production-code fix: load details on the template object already supplied to
importVM, rather than look it up again. Template CPU-detail inheritance is preserved, and explicit VM CPU settings continue to override template defaults. Six regression tests cover the removed dummy template, active-template CPU inheritance, non-VO templates, explicit CPU model/mode overrides, and rejection of a genuinely missing template.Types of changes
Feature/Enhancement Scale or Bug Severity
Feature/Enhancement Scale
Bug Severity
Screenshots (if appropriate)
The acceptance report attached to #13656 includes representative cold VDDK and
warm CBT migration screenshots:
https://github.com/user-attachments/files/32426132/PR13656-acceptance-report.pdf
How Has This Been Tested?
The original change completed its full functional acceptance run:
Blueorangutan smoke testing also passed 156/156 tests:
#13656 (comment)
The original acceptance and smoke results document the restored implementation, but predate the focused fixes in
a5578d6f53b51e2b67c7c537d3b7217c6d44eb3dand6fa7d30b47. For the checkpoint fix, four targeted unit tests, the 38-module Maven package build, and Checkstyle passed. For the cutover failure handling, the 26-module server test reactor and Checkstyle passed (33 targeted CBT tests, including two new tests). Live multi-cycle CBT regression and functional cutover tests have not yet been run on the updated commits. Normal CI and smoke tests should run for this PR.For the dummy-template port in
2a3598ad6a, all sixUserVmImportTemplateTestregression tests passed, and the 26-module server test reactor completed with zero Checkstyle violations. The local command wasmvn -B -pl server -am -Dtest=UserVmImportTemplateTest -Dsurefire.failIfNoSpecifiedTests=false -Dexec.skip=true test. The unrelated schema shell test was skipped viaexec.skipbecause Windows CRLF line endings prevent that script from running locally; this is not a claim that the full test suite passed. Fresh functional acceptance testing is still in progress on the QA environment with the same production fix hot-patched into both management servers. The historical acceptance report above must not be read as a completed acceptance run for this new commit.How did you try to break this feature and the system with this change?
The original acceptance run covered cold and warm migration to NFS, Ceph/RBD
and Linstor, cancellation, retries, invalid state transitions, ownership,
network validation, existing-volume adoption, Windows Server migration,
cleanup and backend leak checks. Full details and evidence are in #13656 and
the linked acceptance report.